Papers with micro F1 score
RoDia: A New Dataset for Romanian Dialect Identification from Speech (2024.findings-naacl)
Copied to clipboard
| Challenge: | a dataset for Romanian dialect identification from speech is released . the dataset includes speech samples from five distinct regions of Romania . |
| Approach: | They propose a dataset for Romanian dialect identification from speech . they propose competitive models to be used as baselines for future research . |
| Outcome: | The first dataset for Romanian dialect identification from speech is released . the top scoring model achieves 59.83% and 62.08%, respectively . |
Misspelling Semantics in Thai (2022.lrec-1)
Copied to clipboard
| Challenge: | In English, more than 70% of documents on the internet contain some form of misspelling . misspellers can be used as prosody to provide additional clues about the writer's attitude . |
| Approach: | They propose two ways to incorporate misspelling semantics into user-generated content . they propose a method to boost micro F1 score by 0.4-2% . |
| Outcome: | The proposed methods can boost the micro F1 score up to 0.4-2% while normalising misspelling is harmful and suboptimal. |